Goto

Collaborating Authors

 true intention


Is Sarcasm Detection A Step-by-Step Reasoning Process in Large Language Models?

arXiv.org Artificial Intelligence

Elaborating a series of intermediate reasoning steps significantly improves the ability of large language models (LLMs) to solve complex problems, as such steps would evoke LLMs to think sequentially. However, human sarcasm understanding is often considered an intuitive and holistic cognitive process, in which various linguistic, contextual, and emotional cues are integrated to form a comprehensive understanding of the speaker's true intention, which is argued not be limited to a step-by-step reasoning process. To verify this argument, we introduce a new prompting framework called SarcasmCue, which contains four prompting strategies, $viz.$ chain of contradiction (CoC), graph of cues (GoC), bagging of cues (BoC) and tensor of cues (ToC), which elicits LLMs to detect human sarcasm by considering sequential and non-sequential prompting methods. Through a comprehensive empirical comparison on four benchmarking datasets, we show that the proposed four prompting methods outperforms standard IO prompting, CoT and ToT with a considerable margin, and non-sequential prompting generally outperforms sequential prompting.


Is the Pope Catholic? Yes, the Pope is Catholic. Generative Evaluation of Non-Literal Intent Resolution in LLMs

arXiv.org Artificial Intelligence

Humans often express their communicative intents indirectly or non-literally, which requires their interlocutors -- human or AI -- to understand beyond the literal meaning of words. While most existing work has focused on discriminative evaluations, we present a new approach to generatively evaluate large language models' (LLMs') intention understanding by examining their responses to non-literal utterances. Ideally, an LLM should respond in line with the true intention of a non-literal utterance, not its literal interpretation. Our findings show that LLMs struggle to generate pragmatically relevant responses to non-literal language, achieving only 50-55% accuracy on average. While explicitly providing oracle intentions significantly improves performance (e.g., 75% for Mistral-Instruct), this still indicates challenges in leveraging given intentions to produce appropriate responses. Using chain-of-thought to make models spell out intentions yields much smaller gains (60% for Mistral-Instruct). These findings suggest that LLMs are not yet effective pragmatic interlocutors, highlighting the need for better approaches for modeling intentions and utilizing them for pragmatic generation.


The Pentagon Wants AI To Reveal Adversaries' True Intentions

#artificialintelligence

From eastern Europe to southern Iraq, the U.S. military faces a difficult problem: Adversaries pretending to be something they're not -- think Russia's "little green men" in Ukraine. But a new program from the Defense Advanced Research Projects Agency seeks to apply artificial intelligence to detect and understand how adversaries are using sneaky tactics to create chaos, undermine governments, spread foreign influence and sow discord. This activity, hostile action that falls short of -- but often precedes -- violence, is sometimes referred to as gray zone warfare, the'zone' being a sort of liminal state in between peace and war. The actors that work in it are difficult to identify and their aims hard to predict, by design. "We're looking at the problem from two perspectives: Trying to determine what the adversary is trying to do, his intent; and once we understand that or have a better understanding of it, then identify how he's going to carry out his plans -- what the timing will be, and what actors will be used," said DARPA program manager Fotis Barlos.


QNX builds in-car speech framework with AT&T's Watson, knows our true intentions

AITopics Original Links

QNX wants to put an end to in-car voice systems that require an awkward-sounding syntax to get the job done. As part of its CES launches, it's rolling out a framework for its speech recognition technology leaning on AT&T's Watson engine. By offloading the phrase interpretation to AT&T's servers, any infotainment system with the framework inside can focus on deciphering the speaker's intent -- letting drivers spend more time navigating or playing music, instead of remembering the necessary magic words. QNX will roll out the voice element as part of its CAR platform at an unspecified point in 2013. We'll have to wait until car and head-end unit designers implement the platform in tangible hardware, but the new speech system will hopefully lead to more organic-sounding conversations with our cars.